Skip to content

feat(agentx): bump Kimi-K3 FP4 MI355X ATOM image to 0924 and track recipe - #3407

Merged
seungrokj merged 10 commits into
mainfrom
amd/kimik3-atom-agentx-0924
Sep 29, 2026
Merged

seungrokj merged 10 commits into
mainfrom
amd/kimik3-atom-agentx-0924

Conversation

@gbyu-amd

@gbyu-amd gbyu-amd commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Track the checked-in ATOM recipe recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382, on image kimi_k3_agentic_0924.

The published concurrency set [1, 4, 14, 16, 48, 56, 72] is unchanged; no points are added or dropped.

  • ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 on every band. ATOM defaults it to 0, so the recipe's prefill attention path was not reached before this change.
  • From concurrency 16 up, ATOM_PREFILL_DECODE_INTERVAL=4 and ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000 hold a ready prefill for four decode passes instead of interleaving it into every step. This is a threshold, not a band: concurrency 1, 4 and 14 run without it.

max-num-seqs, max-num-batched-tokens, gpu-memory-utilization, the CUDA-graph ladder, dcp-size, draft depth, synthetic acceptance, ReplaySSM placement, AITER_REUSE_IDENTICAL_COMM_GROUPS and LMCache sizing are unchanged from #3207.

Re-created on an in-repo amd/ branch so sweep dispatch and labels (AMD, agentx, full-sweep-enabled) apply.

AI model disclosure

Prepared with Claude Code using claude-opus-5 (recipe reconciliation, edits, changelog entry). No other model contributed.

Port to native srt-slurm and LMCache

Merged main in; single-node AgentX now runs on the native recipe inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/kimik3/atom/mi355x-fp4-mtp/agentic.yaml, and the legacy script is gone.

  • The recipe carries this PR's image (rocm/atom-dev:nightly_202609251613), ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1, and the PrefillDelayer from concurrency 16 up.
  • The DCP8 LMCache bands dropped by the native port are back, with the settings from the legacy config and recipes/Agentic-Kimi-K3.md: concurrency 14 and 16 (DSpark 3, ReplaySSM) and 48, 56 and 72 (no draft) on ATOM's in-process lmcache_offload connector, 128 GB/rank up to 48 and 192 GB/rank at 56 and 72, chunk size 1024, PYTHONHASHSEED=0. Concurrency 1 and 4 are unchanged.
  • LMCache on an aggregate ATOM worker uses roles.agg.args.extra-kv-connectors from srt-slurm patch 507-lmcache-server-atom-sglang.patch (feat: support the LMCache server on ATOM and SGLang srt-slurm#32).
  • The single-node adapter reads ATOM's decode-context-parallel-size for the dcp-size check, as it already does for vLLM.

…cipe

Track recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382:
enable FlyDSL FP8 prefill attention on every band, and hold a ready
prefill for four decode passes from concurrency 16 up. The published
concurrency set and every other launch argument are unchanged.

将 MI355X Kimi-K3 FP4 ATOM AgentX 提交切换到 0924 镜像,并跟随
ROCm/ATOM#2382 重调后的 recipe:全部并发开启 FlyDSL FP8 prefill
attention;并发 16 及以上时让就绪的 prefill 等待 4 个 decode 轮次。
已发布的并发点集合与其余启动参数保持不变。

Co-Authored-By: Claude Opus 5 <[email protected]>
@gbyu-amd
gbyu-amd requested a review from a team September 24, 2026 07:44
@gbyu-amd gbyu-amd added AMD agentx AgentX benchmarks, recipes, and infrastructure full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures labels Sep 24, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

将 perf-changelog 条目的 pr-link 指向 PR 3407。

Co-Authored-By: Claude Opus 5 <[email protected]>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Comment thread perf-changelog.yaml Outdated
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Comment thread configs/amd-master.yaml Outdated
@functionstackx

Copy link
Copy Markdown
Collaborator

InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest main into this PR: we have migrated single-node AgentX onto native srt-slurm (#3428), so AgentX configs are now declarative YAML recipes, not per-config 1000+ line bash slop scripts. Please also delete the old benchmarks/single_node/** scripts (see this recipe for the new format).

@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

…x-0924

Port the ATOM image bump and FlyDSL FP8 prefill attention onto the native
srt-slurm recipe; the legacy script is deleted on main.
Carry srt-slurm patch 507 so ATOM aggregate workers accept
extra-kv-connectors, and restore the DCP8 bands from the legacy config
and ROCm/ATOM recipes/Agentic-Kimi-K3.md: concurrency 14 and 16 with
DSpark 3 and ReplaySSM, 48, 56 and 72 without a draft, all on the
in-process lmcache_offload connector (128 GB/rank, 192 GB/rank at 56 and
72). The PrefillDelayer applies from concurrency 16 up.

The single-node adapter now reads ATOM's decode-context-parallel-size
for the DCP_SIZE check, as it does for vLLM.
@cquil11 cquil11 added full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures and removed full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures labels Sep 28, 2026
@seungrokj

Copy link
Copy Markdown
Collaborator

@seungrokj

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 36493779111

seungrokj and others added 2 commits September 29, 2026 09:41
…efill delay, DCP8 LMCache

- Move kimik3-fp4-mi355x-atom-agentic-mtp to rocm/atom-dev:nightly_202609251613,
  tracking recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382.
- Enable FlyDSL FP8 prefill attention (ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1) at
  every concurrency.
- From conc 16 up, hold a ready prefill for four decode passes
  (ATOM_PREFILL_DECODE_INTERVAL=4, ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000);
  conc 1, 4 and 14 unchanged.
- Restore DCP8 LMCache bands on the native srt-slurm recipe: conc 14/16
  (DSpark 3, ReplaySSM) and 48/56/72 (no draft) use ATOM's in-process
  lmcache_offload connector via roles.agg.args.extra-kv-connectors (srt-slurm
  patch 507), 128 GB/rank up to 48 and 192 GB/rank at 56/72. Conc 1 and 4
  stay GPU-resident.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@seungrokj
seungrokj self-requested a review September 29, 2026 17:13

@seungrokj seungrokj left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • insert any additional info here
  • recipe at https://github.com/ROCm/ATOM/blob/main/recipes/Agentic-Kimi-K3.md
  • No change to the Inferact/Kimi-K3-DSpark draft's precision: online_quant_config still excludes every draft linear (layers.*, context_proj), so its weights and activations stay BF16, and it keeps the target's FP8 KV cache (kv_cache_dtype fp8). FlyDSL FP8 prefill attention applies only to the target, since the draft runs its block pass as decode attention.

Signed: seungrokj

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

❌❌❌ REJECTED ❌❌❌

@seungrokj One thing blocks this: the sign-off still has no draft-precision evidence, which the checklist item requires. The additional detail section has not changed since the last verification. Every other check passes. Add the draft checkpoint, its stored precision, how the pinned image handles it by default, and the effective serving precision to the additional detail section, then re-verify.

❌ Check 13 (Draft runs as shipped): FAIL — Draft precision cannot be verified from the sign-off. It says only that the "datatype of the draft model is intact" because the srt-slurm commands did not change. That is not true: this PR bumps the image to rocm/atom-dev:nightly_202609251613, adds DSpark-3 bands at c14/c16 with ReplaySSM and LMCache, and turns on ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 for every band. The sign-off needs to name the draft (Inferact/Kimi-K3-DSpark), give its stored precision (BF16), and say how the pinned image handles it. It should also state that the ptpc_fp8 online_quant_config excludes every draft layer: layers.*.self_attn.{fused_qkv_a,q_b,kv_b,o}_proj, layers.*.mlp.{gate_up,down}_proj and context_proj. Finally, it should say whether FlyDSL FP8 prefill attention touches the draft's attention, and give the effective draft precision. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not set anywhere in the effective recipe (the engine is ATOM).

Passed and not applicable checks

✅ Check 0 (CODEOWNER): PASS — @seungrokj is a named owner of inferencex-e2e/configs/amd-master.yaml. The other changed paths fall only under the * catch-all.

✅ Check 1 (Passing sweep + evals on in-PR commit): PASS — In-PR commit d716d4ca has all 7 agentic / checks (c1, c4, c14, c16, c48, c56, c72) and all 7 agentic eval / checks at success in run 36493779111. Between d716d4ca and pinned head 73874974, the only changes are merges from main. This is an AgentX-only PR, so skipping the fixed-sequence single-node jobs is correct.

✅ Check 2 (Evals pass): PASS — The eval artifacts cover kimi_tool_call_schema_full at all 7 concurrencies: em_strict 0.924–0.963, n_eff 408, infrastructure_success: true. The image is rocm/atom-dev:nightly_202609251613, which matches the PR.

➖ Check 3 (Recipe linked, merged, complete): N/A — This is a single-node ATOM recipe, and the recipe-link requirement only covers single-node vLLM/SGLang recipes. The sign-off links the ATOM recipe for reference.

✅ Check 4 (Reuse command): PASS — seungrokj (COLLABORATOR) posted /reuse-sweep-run 36493779111.

✅ Check 5 (Latest checklist template): PASS — Every item in the current PR_REVIEW_CHECKLIST.md template is present and checked.

✅ Check 6 (Upstream images / engine-first): PASS — (a) Does not apply: the framework is atom. (b) The PR changes the existing ATOM entry kimik3-fp4-mi355x-atom-agentic-mtp rather than adding a new one.

✅ Check 7 (No deprecated models/scenarios): PASS — MODELS.md lists kimik3 agentic coding, with and without DSpark, as active on 2026-09-29.

✅ Check 8 (No architecture hacks): PASS — There are no --hf-overrides flags or model-config edits. FP8 prefill attention and the prefill delayer only change precision or scheduling; they do not remove FLOPs.

✅ Check 9 (Spec-decode via chat template): PASS — The recipe's benchmark env sets AIPERF_APPLY_CHAT_TEMPLATE: 'true'.

✅ Check 10 (No engine patches): PASS — There are no patch files, source-rewriting heredocs or engine wheel installs. The srt-slurm patch 507 changes srtctl, not the engine, and is already on main.

✅ Check 11 (Agentic golden AL): PASS — synthetic_acceptance.py sets ATOM's spec-decode-acceptance-length from kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml (thinking_on). That gives 3.84 for 7 draft tokens (c1/c4) and 3.00 for 3 tokens (c14/c16). c48/56/72 run without a draft.

➖ Check 12 (Append-only): N/A — The new changelog entry does not set append-only: true.

✅ Check 14 (Pareto coverage): PASS — One curve: kimik3 / agentic-coding / MI355X ATOM fp4 / rocm/atom-dev:nightly_202609251613, using the bmk_agentic_* artifacts from run 36493779111 (source SHA d716d4ca). pareto_coverage.py counts 6 of 7 points on the frontier for P90 E2EL and 6 of 7 for P75 E2EL. Total throughput per GPU ranges from 1,619 to 14,526 tok/s.

Assessed commit: 738749748727832a9bbd430e41359dc87fb3f0d8.

@seungrokj
seungrokj self-requested a review September 29, 2026 17:59
seungrokj and others added 2 commits September 29, 2026 11:00
…efill delay, DCP8 LMCache

- Move kimik3-fp4-mi355x-atom-agentic-mtp to rocm/atom-dev:nightly_202609251613,
  tracking recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382.
- Enable FlyDSL FP8 prefill attention (ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1) at
  every concurrency.
- From conc 16 up, hold a ready prefill for four decode passes
  (ATOM_PREFILL_DECODE_INTERVAL=4, ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000);
  conc 1, 4 and 14 unchanged.
- Restore DCP8 LMCache bands on the native srt-slurm recipe: conc 14/16
  (DSpark 3, ReplaySSM) and 48/56/72 (no draft) use ATOM's in-process
  lmcache_offload connector via roles.agg.args.extra-kv-connectors (srt-slurm
  patch 507), 128 GB/rank up to 48 and 192 GB/rank at 56/72. Conc 1 and 4
  stay GPU-resident.
- Draft model precision unchanged: online_quant_config still excludes every
  Inferact/Kimi-K3-DSpark linear (layers.*, context_proj), so its weights and
  activations stay BF16, and it shares the target's FP8 KV cache
  (kv_cache_dtype fp8). FlyDSL FP8 prefill attention applies only to the
  target; the draft's block pass runs as decode attention.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

@seungrokj seungrokj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • insert any additional info here
  • recipe at https://github.com/ROCm/ATOM/blob/main/recipes/Agentic-Kimi-K3.md
  • No change to the Inferact/Kimi-K3-DSpark draft's precision: online_quant_config still excludes every draft linear (layers.*, context_proj), so its weights and activations stay BF16, and it keeps the target's FP8 KV cache (kv_cache_dtype fp8). FlyDSL FP8 prefill attention applies only to the target, since the draft runs its block pass as decode attention.

Signed: seungrokj

@github-actions

Copy link
Copy Markdown
Contributor

✅✅✅ Verdict: PASS ✅✅✅

Passed and not applicable checks

✅ Check 0 (CODEOWNER): PASS — @seungrokj is a listed owner of inferencex-e2e/configs/amd-master.yaml; the other changed paths fall only under the * catch-all.

✅ Check 1 (Passing sweep on in-PR commit): PASS — in-PR commit d716d4ca has all 7 agentic / and 7 agentic eval / check-runs green (c1, 4, 14, 16, 48, 56, 72) in run 36493779111. The Kimi-K3 recipe, master-config entry and ATOM adapter are unchanged between that commit and the pinned head 626b011c, which is still the PR tip. The only later changes are main merges and the changelog.

✅ Check 2 (Evals): PASS — eval_results_all in run 36493779111 has kimi_tool_call_schema_full em_strict 0.924–0.963 (n=408) for all 7 concurrencies, above the 0.0 threshold. Image is rocm/atom-dev:nightly_202609251613, matching the config.

➖ Check 3 (Recipe link): N/A — this is an ATOM single-node recipe, not a vLLM/SGLang one, so the vLLM-recipes/SGLang-cookbook requirement does not apply. For reference, the linked ROCm/ATOM#2382 is MERGED (2026-09-24), and its FlyDSL FP8 prefill and PrefillDelayer (C≥16) settings match.

✅ Check 4 (Reuse command): PASS — /reuse-sweep-run 36493779111 was posted by seungrokj (COLLABORATOR).

✅ Check 5 (Latest checklist template): PASS — every current-template item is present and checked.

✅ Check 6 (Upstream images / engine-first): PASS — the only changed entry, kimik3-fp4-mi355x-atom-agentic-mtp, is an existing ATOM key and this PR does not add a new framework entry. The vLLM Kimi-K3 MI355X submission (#2508) predates the ATOM one (#2697); it was later dropped only when the bash-based scripts were removed in #3463.

✅ Check 7 (Deprecated models): PASS — Kimi-K3 agentic coding with DSpark is an active scenario in MODELS.md.

✅ Check 8 (No architecture hacks): PASS — no --hf-overrides, model-config edits or layer/expert pruning. The changes are image, env toggles, DCP8/LMCache bands and graph sizes.

✅ Check 9 (Spec-decode chat template): PASS — the recipe sets AIPERF_APPLY_CHAT_TEMPLATE: 'true' for AgentX replay.

✅ Check 10 (No engine patches): PASS — no .patch, inline source rewrites, monkey-patching or wheel swaps. setup_script only runs hf download Inferact/Kimi-K3-DSpark, and srt-slurm patch 507 is pre-existing orchestrator (srtctl) plumbing, not the engine.

✅ Check 11 (Agentic golden AL): PASS — the ATOM adapter injects spec-decode-acceptance-length from kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml. Run logs confirm 3.84 for 7 draft tokens (c1, c4) and 3.00 for 3 draft tokens (c14, c16), both matching the thinking_on golden values. c48, c56 and c72 run with no draft.

➖ Check 12 (Append-only): N/A — the new changelog entry is not append-only: true.

✅ Check 13 (Draft as shipped): PASS — the Inferact/Kimi-K3-DSpark safetensors are all BF16 with no quant config. In ATOM's kimi_k3_dspark.py and fnmatch exclude semantics (checked at the ATOM commit nearest the image date), every draft linear (layers.* attention/MLP projections, context_proj) is excluded from ptpc_fp8. markov_head is nn.Embedding and confidence_head is skipped. The draft block pass sets is_prefill=False, so the FlyDSL FP8 prefill path is never reached. Draft KV inherits the target's FP8 pool, which is allowed. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is absent (ATOM engine).

✅ Check 14 (Pareto coverage): PASS — run 36493779111 bmk_agentic_* artifacts (image nightly_202609251613, source d716d4ca, tput_per_gpu vs E2EL) give 6/5 frontier points at P90 and 6/5 at P75, verified with pareto_coverage. c1 is dominated by c4.

Assessed commit: 626b011cdf5ed797f6ded517c409f510807dc477.

@cquil11 cquil11 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Full sweep 36493779111 green.

@seungrokj
seungrokj merged commit d833356 into main Sep 29, 2026
29 checks passed
@seungrokj
seungrokj deleted the amd/kimik3-atom-agentx-0924 branch September 29, 2026 18:17
@cquil11

cquil11 commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 36493779111

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure AMD full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants